Papers with evaluation criteria
Proceedings of the 27th International Conference on Computational Linguistics (C18-1)
Copied to clipboard
| Challenge: | COLING 2018 is the 27th International Conference on Computational Linguistics in Santa Fe, New Mexico. |
| Approach: | the 27th International Conference on Computational Linguistics will be held in Santa Fe, new-mexico . the last COLING in the U.S.A. was held in 1998, as a joint COLing-ACL conference . organizers thank local organizers, program chairs, tutors, volunteers and sponsors . |
| Outcome: | the 27th International Conference on Computational Linguistics will be held in Santa Fe, new-mexico . the last COLING in the U.S.A. was held in 1984, and the only other one was in 1965 . organizers thank the local Organizing committee for their efforts . |
Is He Extroverted? Identifying Missing Relevant Personas for Faithful User Simulation (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing user simulation approaches focus on generating user-like responses in dialogue without verifying whether critical personas are supplied. |
| Approach: | They propose a task of identifying persona dimensions that are relevant but missing in simulating a user's reply for a given dialogue context. |
| Outcome: | The proposed model identifies persona dimensions that are relevant but missing in simulating a user’s response for a given dialogue context. |
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models (2024.naacl-short)
Copied to clipboard
| Challenge: | Existing methods for counter narrative evaluation lack alignment with human judgment as they rely on superficial reference comparisons instead of incorporating key aspects of counter narrative quality as evaluation criteria. |
| Approach: | They propose to use 5 defined aspects to generate counter narrative candidates using human-annotated scores and feedback from counter narrative specialized NGOs to assess their effectiveness. |
| Outcome: | The proposed evaluation framework outperforms existing metrics and achieves strong alignment to human-annotated scores and feedback. |
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Comparative Policy Optimization (CPO) redefines the reward evaluation paradigm by shifting from sample-wise scoring to comparative group-wise score. |
| Approach: | They propose a method to optimize subjective tasks by shifting from sample-wise to comparative group-wise scoring. |
| Outcome: | The proposed framework shifts from sample-wise scoring to comparative group-wise score . it minimizes contextual bias and enables more robust and fair performance evaluation. |
Inject Rubrics into Short Answer Grading System (D19-61)
Copied to clipboard
| Challenge: | Short Answer Grading (SAG) is a task of scoring students’ answers in examinations. Existing SAG systems only predict scores based on the answers, but they ignore important evaluation criteria such as rubrics. |
| Approach: | They propose to inject rubrics into SAG models by introducing word-level attention mechanism into the model to locate information in each answer that are highly related to the score. |
| Outcome: | The proposed model outperforms the state-of-the-art model on the widely used ASAP-SAS dataset under low-resource settings. |
ZhuJiu-Knowledge: A Fairer Platform for Evaluating Multiple Knowledge Types in Large Language Models (2024.naacl-demo)
Copied to clipboard
| Challenge: | evaluating the knowledge of large language models (LLMs) is crucial, and rapid advancement in large language modeling has heightened the importance of model evaluations. |
| Approach: | They propose a fairer benchmark for evaluating multiple knowledge types of LLMs by focusing on commonsense knowledge, world knowledge, and language knowledge. |
| Outcome: | The proposed framework evaluates 14 current mainstream LLMs and provides a detailed discussion and analysis of their results. |
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge (2026.acl-demo)
Copied to clipboard
| Challenge: | AutoChecklist is an open-source library that unifies checklist-based evaluation into composable pipelines. |
| Approach: | They propose an open-source library that unifies checklist-based evaluation into composable pipelines. |
| Outcome: | The open-source library unifies checklist-based evaluation into composable pipelines. |
JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for detecting factual hallucinations in generated content exhibit limitations in the first two stages of the halluciation detection pipeline. |
| Approach: | They propose a joint claim-and-query generation framework that can detect factual hallucinations in generated content. |
| Outcome: | The proposed method outperforms existing methods on open-domain QA hallucination detection benchmarks. |
A Practical Incremental Learning Framework For Sparse Entity Extraction (C18-1)
Copied to clipboard
| Challenge: | Existing approaches to extract entities from textual data are expensive and unattractive due to the high cost of training. |
| Approach: | They propose a framework that integrates Entity Set Expansion and Active Learning to reduce the cost of data annotation. |
| Outcome: | The proposed framework reduces the cost of sparse entity annotation by 85% and 45% while maintaining high accuracy. |
Evaluating Step-by-step Reasoning Traces: A Survey (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation practices are inconsistent, resulting in fragmented progress across evaluator design and benchmark development. |
| Approach: | a survey provides a comprehensive overview of step-by-step reasoning evaluation . existing evaluation practices are inconsistent, resulting in fragmented progress . |
| Outcome: | The proposed evaluation criteria are based on four top-level categories . the results are presented in a systematic review of the literature. |
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Question answering (QA) tasks have been extensively studied in the field of natural language processing. |
| Approach: | They propose a method that leverages large language models and the analytic hierarchy process to assess open-ended questions. |
| Outcome: | The proposed method more closely aligns with human judgment compared to baselines on four datasets. |
Rogue Scores (2023.acl-long)
Copied to clipboard
| Challenge: | Critical evaluation decisions and parameters are routinely omitted, making most reports irreproducible . Thousands of papers use nonstandard evaluation packages with software defects that produce incorrect scores. |
| Approach: | a systematic review of over two thousand papers using a popular metric called ROUGE finds errors . critical evaluation decisions and parameters are routinely omitted, making most reported scores irreproducible . a large number of ROUGEE model evaluation scores have been incorrectly computed . |
| Outcome: | a systematic review of over two thousand papers finds that ROUGE scores are incorrect . the metric is widely used in machine learning and is inconsistent with human evaluations . |
Visual Question Decomposition on Multimodal Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for question decomposition focus on unimodal language models, but question decomposing capability of Multimodal Large Language Models (MLLMs) has yet to be explored. |
| Approach: | They propose a finetuning dataset and a training objective for selective decomposition to enhance the model's question decomposing capability. |
| Outcome: | The proposed dataset shows that existing models struggle to produce high-quality sub-questions. |
CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling (2025.findings-acl)
Copied to clipboard
Taneesh Gupta, Shivam Shandilya, Xuchao Zhang, Rahul Madhavan, Supriyo Ghosh, Chetan Bansal, Huaxiu Yao, Saravan Rajmohan
| Challenge: | Reward modeling in large language models is susceptible to reward hacking . flawed reward signals often lead to outputs that optimize for spurious correlates . |
| Approach: | They propose a new approach that generates dynamic, context-relevant criteria to ground the reward model prior to producing reward scores. |
| Outcome: | The proposed approach generates dynamic, context-relevant criteria to ground the model prior to producing reward scores. |
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices (2025.emnlp-industry)
Copied to clipboard
Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Ridwan Mahbub, Mizanur Rahman, Amran Bhuiyan, Israt Jahan, Mir Tafseer Nayeem, Shafiq Joty, Enamul Hoque, Jimmy Huang
| Challenge: | Large Vision-Language Models (LVLMs) with only 7B parameters perform poorly as judges in resource-constrained settings. |
| Approach: | They propose two approaches to ensure costefficient evaluation by combining multiple criteria into a single query and domainadaptive transfer learning to create a 2Bparameter VLM on a chart dataset. |
| Outcome: | The proposed model can effectively transfer knowledge from one dataset to another to make it a more specialized model. |
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies show that doctors can save significant amounts of time when using automatic note generation. |
| Approach: | They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics. |
| Outcome: | The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets. |
DataSciBench: An LLM Agent Benchmark for Data Science (2026.findings-acl)
Copied to clipboard
Dan Zhang, Sining Zhoubian, Min Cai, Fengzu Li, Lekang Yang, Wei Wang, Tianjiao Dong, Ziniu Hu, Jie Tang, Yisong Yue
| Challenge: | Existing benchmarks focus on single task, simple evaluation metrics, and readily available ground truth (GT) DataSciBench is built on curated, natural, and challenging prompts with complex evaluation criteria and uncertain GT. |
| Approach: | They propose a benchmark for evaluating Large Language Models in data science that integrates LLM-based self-consistency and human verification to ensure accuracy. |
| Outcome: | The proposed framework outperforms open-source models in all metrics and offers rigorous insights into LLM strengths and weaknesses. |
Building the Directed Semantic Graph for Coherent Long Text Generation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for conditional long text generation ignore the coherence issue of the generated texts. |
| Approach: | They propose a two-stage approach to generate coherent long text based on short input text . they first build a document-level path for each output text with each sentence embedding as its node . |
| Outcome: | The proposed approach is superior to state-of-the-art approaches on three real-world datasets. |
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models (2024.emnlp-main)
Copied to clipboard
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | Existing open-source evaluation paradigms lack flexibility and performance . language model-based evaluation is cheap and scalable, but it is difficult to evaluate . |
| Approach: | They propose a language model-based evaluation paradigm that uses a scalar indicator of quality to assess LM outputs. |
| Outcome: | The proposed language model-based evaluation model is more powerful than its predecessor. |
Question Answering in Climate Adaptation for Agriculture: Model Development and Evaluation with Expert Feedback (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing domain-specific question answering systems have generative capabilities, but their ability to answer climate adaptation questions remains unclear. |
| Approach: | They propose an iterative framework that enables LLMs to dynamically aggregate information from heterogeneous sources, such as climate literature and structured tabular climate data from climate model projections and historical observations. |
| Outcome: | The proposed framework enables LLMs to dynamically aggregate information from heterogeneous sources, such as text from climate literature and structured tabular climate data from climate model projections and historical observations. |
Logic Traps in Evaluating Attribution Scores (2022.acl-long)
Copied to clipboard
| Challenge: | Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. |
| Approach: | They propose to review existing methods for evaluating attribution scores and summarize the logic traps in these methods. |
| Outcome: | The proposed methods show that they do not contain logic traps and that they are not reliable. |
Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework (2025.findings-acl)
Copied to clipboard
Kaishuai Xu, Tiezheng Yu, Yi Cheng, Wenjun Hou, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li
| Challenge: | Existing methods for fine-tuning open-source LLMs are limited to text-based analysis under predefined general criteria. |
| Approach: | They propose a framework that fine-tunes LLMs to replicate the evaluation explanations and judgments of proprietary models. |
| Outcome: | The proposed evaluation framework outperforms existing fine-tuned evaluation methods in effectiveness and robustness. |
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability. |
| Approach: | They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria. |
| Outcome: | The proposed system is based on 11 common aspects with different evaluation criteria. |
Reward Modeling for Scientific Writing Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models for scientific writing evaluation are primarily optimized for general-purpose benchmarks with fixed scoring rubrics and evaluation criteria. |
| Approach: | They propose to train scientific writing evaluation models that leverage domain knowledge . they use a two-stage evaluation framework that optimizes evaluation preferences and refines reasoning capabilities . |
| Outcome: | The proposed model generalizes effectively across tasks and to previously unseen settings. |
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs. |
| Approach: | They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria. |
| Outcome: | The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning. |
What is the Best Way for ChatGPT to Translate Poetry? (2024.acl-long)
Copied to clipboard
| Challenge: | Despite promising results, our analysis reveals persistent issues in the translations generated by ChatGPT that warrant attention. |
| Approach: | They propose an Explanation-Assisted Poetry Machine Translation method which leverages monolingual poetry explanation as a guiding information for the translation process. |
| Outcome: | The proposed method outperforms traditional translation methods of ChatGPT and the existing online systems in English-Chinese poetry translation. |
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models (2026.findings-acl)
Copied to clipboard
Qingjia Huang, Jingyu Zhang, Jianguo Wu, Yakai Li, Weijuan Zhang, Yankai Rong, Junyi Yao, Shengzhi Zhang, Xiaoqi Jia
| Challenge: | Currently, evaluation criteria and methods used for jailbreak effectiveness are inconsistent. |
| Approach: | They propose a framework to measure jailbreak effectiveness using a model that filters out jailbreak noise while preserving the original malicious question. |
| Outcome: | The proposed framework outperforms existing evaluation methods on a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. |
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation protocols for text generation suffer from rating inconsistencies . lexical overlap-based metrics align poorly with human judgments . |
| Approach: | They propose a checklist-based evaluation framework that improves rating reliability via decomposed binary questions. |
| Outcome: | The proposed framework improves rating reliability by decomposing binary questions . it improves agreement across evaluator models by 0.45 and reduces score variance . human evaluation remains the gold standard, but it #, Equal contribution. |
CE-RM: A Pointwise Generative Reward Model Optimized via Two-Stage Rollout and Unified Criteria (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have shown that rule-based evaluation methods are ineffective for open-ended natural language generation. |
| Approach: | They propose a pointwise generative reward model with a dedicated two-stage rollout method and unified query-based criteria that can be trained with 5.7K high-quality data. |
| Outcome: | The proposed model achieves superior performance on diverse reward model benchmarks, especially in Best-of-N scenarios, and delivers more effective improvements in downstream RL practice. |
Substance over Style: Evaluating Proactive Conversational Coaching Agents (2025.acl-long)
Copied to clipboard
Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush, Isaac Galatzer-Levy, Shwetak Patel, Daniel McDuff, Tim Althoff
| Challenge: | Recent NLP research has focused on single-turn tasks with well-defined objectives or evaluation criteria. |
| Approach: | They describe five multi-turn coaching agents that exhibit distinct conversational styles and evaluate them through a user study. |
| Outcome: | The authors compare user feedback with third-person evaluations from health experts and an LM to find that stylistic components in absence of core functionality are viewed negatively. |
T2R-BENCH: A Benchmark for Real World Table-to-Report Task (2025.emnlp-main)
Copied to clipboard
Jie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei, Yu Zhao, Xiangyu Li, Jiaxin Peng, Xiaoyan Gu, Jian Yang, Wenhan Chang, Zhenhe Wu, Jiang Zhong, Shuangyong Song, Xuelong Li
| Challenge: | Existing table benchmarks lack the capacity to adequately assess the practical application of table reasoning in industrial applications. |
| Approach: | They propose a bilingual table-to-report task and a table-based benchmark to assess the quality of table reasoning. |
| Outcome: | The proposed task is based on a bilingual benchmark with 457 industrial tables and evaluation criteria to measure the quality of report generation. |
From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation (2026.findings-acl)
Copied to clipboard
Pujun Zheng, Jiacheng Yao, Jinquan Zheng, Chenyang Gu, Guoxiu He, Jiawei Liu, Yong Huang, Tianrui Guo, Wei Lu
| Challenge: | Large language models (LLMs) are currently used to evaluate scientific papers by assigning an absolute score to each paper independently. |
| Approach: | They propose a comparison-native framework for paper evaluation that integrates comparison into both data construction and model learning. |
| Outcome: | The proposed framework achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B, while exhibiting robust generalization to five previously unseen datasets. |
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators (2025.findings-acl)
Copied to clipboard
Bosi Wen, Pei Ke, Yufei Sun, Cunxiang Wang, Xiaotao Gu, Jinfeng Zhou, Jie Tang, Hongning Wang, Minlie Huang
| Challenge: | Existing efforts to optimize text evaluation prompts neglect the combinatorial impact of multiple factors, leading to insufficient optimization of the evaluation pipeline. |
| Approach: | They propose to integrate 8 key factors for evaluation prompts and integrate them into an algorithm that searches for well-behaved prompting strategies for LLM evaluators. |
| Outcome: | The proposed method outperforms existing methods and human-designed evaluation prompts on four evaluation tasks. |
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)
Copied to clipboard
Junyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma
| Challenge: | Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation. |
| Approach: | They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation. |
| Outcome: | The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics. |
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review (2026.findings-acl)
Copied to clipboard
| Challenge: | a growing workload has made peer review automation an urgent necessity, says a new study . official conference guidelines and reviewer-imitating guidelines degraded review performance . current human-based peer review system faces serious challenges, authors say . |
| Approach: | They analyze how reviewer guidelines influence automated peer review . official conference guidelines produce review results consistent with human judgments . |
| Outcome: | The proposed reviewer guidelines produce results consistent with human judgments . the proposed reviewers' imitations degraded performance, the authors note . |